Back

Cell Systems

Elsevier BV

Preprints posted in the last 90 days, ranked by how well they match Cell Systems's content profile, based on 201 papers previously published here. The average preprint has a 0.20% match score for this journal, so anything above that is already an above-average fit.

1
Reliable single-cell perturbations explain and improve model performance

Wang, X.; Kuipers, J.; Hugi, F.; Platt, R. J.; Beerenwinkel, N.

2026-08-12 bioinformatics 10.64898/2026.08.11.744177 medRxiv
Top 0.1%
51.0%
Show abstract

Predicting single-cell transcriptional responses to perturbations is central to building the virtual cell, yet recent benchmarks show that simple baseline methods often outperform complex models, and model comparisons depend on the evaluation metric. Most studies assume that preprocessed RNA sequencing data are reliable ground truth for both training and evaluation. Here, we test this assumption by measuring the reliability of perturbations and their alignment with shared perturbation responses, classifying each perturbation as specific, shared, or unreliable. Among 7,170 perturbations from 29 datasets, 65% are unreliable, 11% shared, and 24% specific. Applying these quality labels to published benchmarks shows that model comparisons depend on perturbation quality. Training with reliable perturbations alone matches or outperforms full-data performance while using 55% of all training perturbations. Our framework also enables prospective experimental design: for most perturbations, a 28-cell pilot experiment accurately predicts how many cells a full screen needs to be reliable.

2
COMPASS: Component-Wise Inference of Shared and Gene-Specific Perturbation Response

Liang, H.; Singh, R.

2026-08-06 bioinformatics 10.64898/2026.08.03.742643 medRxiv
Top 0.1%
39.4%
Show abstract

Predicting how a genetic perturbation reshapes a cells transcriptome is a central goal of computational biology. Previous studies report that the mean response across training perturbations rivals specialized models on standard accuracy metrics, even though it cannot distinguish which perturbation occurred. Across 2,270 CRISPRi perturbations measured in each of six cell lines, we show that this apparent paradox reflects a conserved organization of perturbation responses. Perturbations span a continuum from responses strongly aligned with the mean to more targeted responses that depart from it. Crucially, a perturbations position along this continuum is conserved across cell lines (Kendalls W = 0.59) and predictable from STRING protein-interaction embeddings (R2 = 0.35). We formalize this structure with COMPASS, an interpretable linear model that decomposes each response into shared and gene-specific components and estimates them separately. The shared-response component is modeled as a cell-line-wide response scaled by a perturbation-specific coefficient. This coefficient is strongly conserved across cell lines. The residual gene-specific component--which is moderately conserved across cell lines--recovers pathway-level programs. COMPASS outperforms scGPT, CPA, GEARS, GenePert, and SO_SCPLOWTATEC_SCPLOW in both response accuracy (de-biased Pearson delta 0.34 vs. [≤] 0.32) and perturbation discrimination (cosine PDS gain 0.23 vs. [≤] 0.08). These results recast perturbation prediction across cellular contexts as component-wise inference, with each component estimated from the evidence best suited to it.

3
Directing hierarchical cell fate decisions through sequential pulses of minimal signaling alphabets

Ji, Y.; Li, Z.

2026-07-29 systems biology 10.64898/2026.07.28.741174 medRxiv
Top 0.1%
33.7%
Show abstract

Inductive signals can direct cell fate decisions, yet the diversity of cell identities far exceeds the available signaling repertoire. While developing systems resolve this paradox by applying temporal sequences of minimalistic signals to hierarchically organized gene regulatory networks (GRNs), the rules governing robust sequential fate navigation remain largely unknown. To address this, we modeled how sequential signals steer trajectories toward arbitrary fates within a Waddington-like manifold, enumerating parameters and combinatorial logics of cross-inhibition and self-activation (CIS) modules, and various ways of applying the sequential signals. We uncover a fundamental conflict between commitment stability and inductive plasticity that severely limits sequential fate navigation. Crucially, introducing non-inductive gap intervals systematically resolves this bottleneck by dissipating kinetic leakage without sacrificing upstream stability. Finally, we demonstrate that polarization-division cycle can autonomously generate signals meeting these temporal requirements. Together, our findings provide a minimal model demonstrating the feasibility and design principles for driving hierarchical cell fate decisions through sequential pulses of a restricted signaling repertoire.

4
Time-resolved operator archetypes characterize dynamical sensitivity during cell-state transitions

Redd, D. M.; Green, S. G.; Terooatea, T. W.

2026-08-23 bioinformatics 10.64898/2026.08.21.745996 medRxiv
Top 0.1%
32.8%
Show abstract

During development, cells traverse gene expression states where their local dynamical sensitivity changes sharply, yet existing computational methods provide limited access to when and where this sensitivity peaks along a trajectory. Here we introduce scJDO (single-cell Jacobian Differential Operators), a framework that characterizes how local dynamical sensitivity evolves during cell fate transitions, together with an explicit account of what that representation can and cannot recover from snapshot data. scJDO treats time-indexed Jacobians as explicit analytical objects, projecting the temporal sequence of operators into a shared subspace and decomposing it into recurrent operator archetypes with interpretable temporal activation profiles. Unlike methods that derive Jacobians from splicing-kinetic vector fields, scJDO learns a neural drift field directly from cell-state geometry via diffusion score matching, enabling Jacobian analysis on trajectory-resolved scRNA-seq datasets regardless of splicing-data availability. Applied to a dense time-course of induced pluripotent stem cell (iPSC) reprogramming, scJDO resolves a quantitative operator-level signature that distinguishes diverted from productive fate: the productive trajectory executes a sequential handoff from an early MEF-exit operator regime to a late pluripotency-associated regime, whereas the diverted trajectory maintains the early regime and instead activates a distinct stress-associated archetype. We validate scJDO across four settings: synthetic benchmarks with analytically known ground truth, branching hematopoiesis, dense real time-course reprogramming, and a perturbational setting using Schrodinger bridges in K562 CRISPRi Perturb-seq. We compare against the two most widely used single-cell Jacobian methods on a dataset where all three are runnable, finding that scJDO shares significantly more gene-level and directional operator structure with Dynamo than expected by chance while providing operator-level analysis on datasets without splicing kinetics. We further characterize the boundary of the representation directly. At a fate-decision saddle, eight mathematically distinct readouts of the same learned drift field are consistent with a single explanation: a drift field fit to snapshot density reproduces density-dominant separation between committed branches rather than the low-variance transverse instability that defines the decision. Together, scJDO provides an operator-level view of single-cell dynamics and an explicit characterization of its own identifiability boundary.

5
Virtual-cell models compress unseen intervention geometry through a target-specific generalization bottleneck

Huang, Y.; Wang, H.; Wilson, P. C.

2026-08-24 bioinformatics 10.64898/2026.08.21.746243 medRxiv
Top 0.1%
31.3%
Show abstract

Predictive models of cellular perturbation are often judged by how closely they reconstruct molecular states after unseen interventions. We show that high state-level similarity can coexist with loss of the relationships that distinguish perturbations, a failure we term Intervention Geometry Compression (IGC). Across established models and perturbation settings, unseen interventions show weakened global and local geometry, reduced between-intervention variance and spectral collapse. The failure is not primarily explained by response-space capacity. Instead, diagnostic projections localize much of the missing geometry to a small number of residual response directions learned from seen interventions; these directions outperform complexity-matched random subspaces and replicate in an independent Jiang perturbation resource. Polarity captures part, but not all, of this continuous orientation signal. Time-resolved analyses further show that correct trajectory entry markedly improves downstream propagation, while a held target's own early empirical response rapidly reveals endpoint orientation. Finally, same-target empirical anchoring transfers intervention identity across contexts far more effectively than increasing exposure to other interventions. These results identify intervention-coordinate assignment as an information bottleneck in virtual-cell generalization and support a design principle: empirically anchor intervention identity, then use models to generalize anchored effects across cellular contexts.

6
Perturbation response decomposition enables biologically aligned generalization to unseen perturbations and cellular contexts

Molina, A.; Zhang, X.

2026-07-27 bioinformatics 10.64898/2026.07.24.740459 medRxiv
Top 0.1%
31.3%
Show abstract

Predicting single-cell responses to genetic perturbations could reveal the vast combinatorial space of perturbations and cellular contexts that is infeasible to measure experimentally, yet current deep learning models generalize poorly and often fail to outperform simple baselines. Here we demonstrate that generalizability in perturbation prediction requires identifying and representing distinct components of cellular response rather than on increasing model complexity alone. We introduce a decomposition framework that explicitly separates transcriptional responses into global, perturbation-specific, cell-line-specific, and perturbation-by-cell-line interaction components. Applied to four CRISPR-interference Perturb-seq screens on multiple cell lines, our framework reveals that these components have distinct structures and information requirements. The global response component is low-dimensional, reflects recurrent proliferation and stress response programs, and can be inferred from control gene expression. In contrast, the perturbation and cell-line specific components are high-dimensional and cannot be recovered from control expression. We therefore develop response-component-aligned models that map biological priors, such as gene coessentiality, onto the geometry of observed transcriptional responses. Critically, this alignment enables simple linear or multilayer perceptron (MLP) based models to outperform state-of-the-art architectures across multiple generalization settings, including unseen cell lines and combinations of unseen perturbations. Together, our framework for response decomposition and alignment provides a principled basis for evaluating and designing perturbation-prediction models, showing that generalization depends primarily on matching biological information to the response components rather than on model complexity alone.

7
A hybrid machine learning and enzyme-constrained metabolic model for ab initio prediction of proteome reallocation

Motamedian, E.; Nikoloski, Z.

2026-07-21 systems biology 10.64898/2026.07.20.739489 medRxiv
Top 0.1%
30.9%
Show abstract

High expression of heterologous proteins in microbial cell factories frequently triggers a severe burden due to reallocation of finite cellular proteome. Conventional constraint-based models struggle to predict these resource shifts ab initio without relying on condition-specific omics data. To bridge this gap, we developed the Hybrid Transcription-Translation (HyTT) framework, combining multivariate adaptive regression splines (MARS) with enzyme-constrained metabolic models by enforcing an 80S ribosome integrity constraint. Cast as a mixed-integer linear programming problem, HyTT mathematically couples macroscopic spatial boundaries with microscopic, sequence-derived translational costs based on a bisection search. Validation against steady-state chemostat quantitative proteomics data demonstrated the superior capability of HyTT over contenders in predicting system-wide resource (re)allocation in Saccharomyces cerevisiae. Operating ab initio, the framework doubled the predictive accuracy of protein abundances (Pearson r=0.501) compared to conventional models, successfully segregating the minimal essential proteome from the cellular reserve pool. Crucially, HyTT autonomously captures complex stress responses vital for metabolic engineering. Upon simulating a 15% recombinant protein burden, the framework accurately predicted systemic growth retardation, decrease of ribosomal portion of the proteome, and surge of ethanol production, in line with the Crabtree effect. System-level analysis uncovered that cells adapt to restricted proteomic capacity through non-uniform metabolic rerouting, downregulating respiratory complexes in favor of high-turnover glycolytic enzymes, and relying on ribosomal paralog switching to minimize sequence-specific assembly costs. Ultimately, HyTT provides a computationally agile, sequence-driven platform for decoding dynamic resource reallocation, offering a powerful predictive tool to navigate metabolic trade-offs and guide rational strain design without requiring condition-specific multi-omics inputs.

8
Perturbation Curve models continuous transcriptional response trajectories and improves prediction of genetic modulations

Zhong, Y.; wang, l.; Yang, G.; Yu, L.; Qi, X.; Jiang, H.

2026-06-19 bioinformatics 10.64898/2026.06.16.732192 medRxiv
Top 0.1%
29.8%
Show abstract

Single-cell CRISPR screens, Perturb-seq, have revolutionized functional genomics by revealing biological causality. However, although perturbation assignments are typically represented as discrete labels, the cell-level effective strength of perturbations is often continuous and diverse. Current analytical frameworks struggle to decouple the variability in perturbation strength from the diversity of downstream responses. Here, we present Perturbation Curve (PertCurve), a nonlinear, curve-based computational framework that models the trajectories of transcriptomic responses by explicitly incorporating diverse perturbation magnitudes and strengths. By ordering cells by perturbation strength, we demonstrate that PertCurve accurately recapitulates the response magnitudes and reveals the distinct modularity and asynchrony patterns of downstream gene behaviors. These patterns are categorized into archetypes, including proportional, sensitive, and threshold responses. By applying this framework across CRISPRi/a modalities, we identify universal response patterns in viral infection, apoptosis, and proliferation genes, and reveal previously overlooked context-specific regulatory features in cell differentiation. Finally, incorporating PertCurve into perturbation prediction models and evaluation metrics enhances predictive performance, delivering actionable insights for refining established models.

9
Scalable biophysical constraints for physiologically consistent metabolic states

Toumpe, I.; Weilandt, D. R.; Narayanan, B.; Fengos, G.; Hatzimanikatis, V.; Miskovic, L.

2026-07-09 systems biology 10.64898/2026.07.03.736321 medRxiv
Top 0.1%
28.0%
Show abstract

Systems biology aims to develop predictive models that connect molecular mechanisms to cellular behavior. Genome-scale metabolic models are among the most widely used frameworks for integrating stoichiometric, thermodynamic, and omics-derived information to predict feasible metabolic phenotypes. However, cellular metabolism operates on timescales governed by enzyme kinetics and by the relationship between metabolic fluxes and metabolite pool sizes. In steady-state metabolic models, this relationship can be expressed in terms of metabolite turnover rates, defined as flux-to-pool-size ratios that quantify how rapidly metabolite pools are renewed. As a result, physiologically consistent steady-state solutions should not only satisfy mass-balance and thermodynamic constraints but also exhibit turnover rates consistent with enzyme-mediated cellular dynamics. Current constraint-based approaches can admit many steady-state flux-concentration states that do not account for turnover rates, resulting in phenotypes incompatible with realistic metabolic dynamics, even when multiple types of data are imposed. Here, we present METEOR-K, an optimization framework that links steady-state metabolic fluxes to metabolite concentrations via turnover rate constraints to identify dynamically plausible flux-concentration reference states. Because these constraints reshape the feasible solution space, we also introduce turnover-rate-aware sampling strategies to efficiently explore the resulting feasible region. We applied METEOR-K to models of increasing scope and scale, including a reduced glycolysis pathway, anaerobic E. coli, and near-genome-scale ovarian cancer models. METEOR-K narrowed the admissible steady-state solution space, reduced uncertainty in feasible flux-concentration states, and improved local dynamic behavior. In nonlinear ODE simulations of bioreactor cultivation and drug-response scenarios, METEOR-K-derived states produced intracellular response times compatible with growth-supporting metabolic operation and perturbation recovery. Overall, these results establish metabolite turnover rates as scalable biophysical constraints that improve the physiological consistency of steady-state metabolic modeling. Because turnover rates encode flux-to-pool-size timescale constraints, METEOR-K moves part of physiological-consistency assessment upstream of kinetic parameterization, yielding better-suited flux-concentration reference states for kinetic modeling and dynamic prediction.

10
Double Machine Learning with Multi-Gene Shared Backgroundfor Causal Inference in Single-Cell Data: Grouping Deviation Follows a Random Walk and the Accuracy-Compute Trade-Off

Ye, W.; Jiang, X.; Shen, F.

2026-08-10 systems biology 10.64898/2026.08.08.743701 medRxiv
Top 0.1%
27.8%
Show abstract

In high-throughput single-cell transcriptomics (p {approx} 20,000 genes), performing double machine learning (DML) causal inference on q {approx} 5,000 target genes requires nuisance function fits that grow linearly with the number of targets (Kf cross-fitting folds, Kf = 5 or 10), far exceeding feasible computational budgets, especially with deep learning. We propose a Randomized Partition Strategy (RPS): randomly divide target genes into groups, share one background compression per group, reducing deep learning model training to q/m runs (m = group size) --- a factor of m savings. The cost of grouping is accuracy loss --- we prove that the cumulative deviation of the estimator follows a one-dimensional drift-free symmetric random walk, with diffusion variance growing linearly with group size and mean squared displacement equaling the mean squared error, so accuracy loss is predictable: m = 1 is always optimal, accuracy cost is monotonically increasing, and a small accuracy sacrifice yields m-fold compute savings. On GSE189050 SLE single-cell data (Memory B cells, n = 2120), both PCA and DL methods converge to the same conclusion, confirming the random walk mechanism is method-independent; an unexpected finding is that DL diffusion growth is only 16%, far slower than PCAs 7.4 times. This work provides a quantifiable theoretical foundation for compute strategy selection in single-cell high-dimensional causal inference.

11
Towards Principled Evaluation of Single-Cell Perturbation Prediction Models

Schäfer, P. S. L.; Reid, K.; Boldyga, Z.; Aksu, E. D.; Hakem, H.; Saez-Rodriguez, J.

2026-07-27 bioinformatics 10.64898/2026.07.23.740433 medRxiv
Top 0.1%
27.3%
Show abstract

Single-cell perturbation experiments measure how interventions alter cellular phenotypes. However, the number of possible perturbations and biological contexts far exceeds what can be tested experimentally. Motivated by this constraint, predictive models aim to extrapolate cellular responses to unseen conditions. Despite substantial efforts in model development, benchmark studies have reached inconsistent conclusions about the capabilities of current perturbation-response models. A major challenge is that evaluation protocols vary widely across studies, making results difficult to compare. Furthermore, the lack of consensus on evaluation hampers progress because it is unclear which predictive capabilities new models should prioritize. To help build consensus on evaluation principles, we develop a taxonomy that decomposes evaluation protocols into their representation, metric, score transformation, and reporting strategies. We characterize how these choices determine which aspects of prediction quality a benchmark measures and discuss criteria for selecting and assessing protocols in relation to specific benchmarking goals. We additionally provide scPertEval, a Python package with reference implementations of selected evaluation protocols, and use it to assess protocol behavior across seven publicly available single-cell perturbation datasets. By making evaluation choices and their underlying trade-offs explicit, we aim to stimulate a community discussion about developing more comparable and task-aligned evaluation protocols. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=129 SRC="FIGDIR/small/740433v1_ufig1.gif" ALT="Figure 1"> View larger version (32K): org.highwire.dtl.DTLVardef@bd1e3corg.highwire.dtl.DTLVardef@c26c2org.highwire.dtl.DTLVardef@1c49afdorg.highwire.dtl.DTLVardef@9b988c_HPS_FORMAT_FIGEXP M_FIG C_FIG

12
Learning protein function through autonomous experimental interaction

Brooks, C.; Notin, P.; Romero, P. A.

2026-08-20 bioengineering 10.64898/2026.08.14.744985 medRxiv
Top 0.1%
27.2%
Show abstract

Biological AI learns primarily from existing observations, but many questions cannot be answered from available data alone. Here we show that AI can instead acquire knowledge by acting directly on biological systems and learning from the consequences. We developed a closed-loop framework in which autonomous agents design protein variants, construct and characterize them in a robotic laboratory, learn from the resulting experimental feedback, and decide what experiments to perform next. We then allowed the system to operate continuously and without human intervention for approximately one month, during which multiple agents independently explored protein sequence space while learning from shared experimental experience. Applied to glycoside hydrolases, the agents discovered enzymes with substantially altered substrate specificity toward non-native sugars and progressively learned the structure of the underlying sequence-function landscape. The resulting experimental experience also revealed determinants of substrate specificity and protein expression that were not specified as learning objectives. These results demonstrate that AI can autonomously interact with biology over extended periods to acquire knowledge through experience, establishing a framework for biological discovery driven by continuous experimental interaction.

13
A permutation-free family-wise error rate for the moderated top-gene scan under gene correlation

Dwyer, W. J.

2026-08-21 bioinformatics 10.64898/2026.08.17.745282 medRxiv
Top 0.1%
26.9%
Show abstract

A differential-expression scan reports the genes with the largest moderated t-statistics, so controlling the family-wise error rate means controlling the null distribution of the maximum statistic over genes. Under gene correlation this is widely believed to require permutation, because correlation changes the effective multiplicity and corrupts the empirical-Bayes variance prior behind the moderated t-statistic. We decompose that liberality by an error-budget ablation and show that, within the simulated model class, it reduces principally to an inflation of the empirical-Bayes prior degrees of freedom: substituting the true prior returns the family-wise error to the independent-gene small-sample baseline, so dependence imposes no separate barrier once the prior is correct. Correlation deflates the cross-gene spread of the log sample variances; because the prior degrees of freedom decreases in that spread, the prior is over-estimated and the moderated maximum turns liberal. Dividing the observed spread by one minus the mean squared gene correlation, estimated by a tuning-free spectral U-statistic with an unbiased trace target, reverses the mechanism and holds the family-wise error near the baseline at retained power without permutation. An observable instability index flags when severe co-expression should defer to permutation.

14
MSGPCA: Multi-Slice Graph PCA for replicate-aware Spatial Omics analysis

Chakraborty, A.; Neelon, B.; Lawson, A.; Angel, P.; Chung, D.; Seal, S.

2026-08-11 bioinformatics 10.64898/2026.08.05.743056 medRxiv
Top 0.1%
26.8%
Show abstract

As spatial transcriptomics (ST) and spatial proteomics (SP) technologies mature, experimental designs are increasingly moving beyond single-slice analyses toward multi-slice studies involving one or more donors and experimental conditions. Although these designs enable the identification of reproducible spatial signals, they also introduce substantial biological heterogeneity, particularly when integrating non-serial slices or anatomically distinct regions. If not modeled carefully, such variation can blur slice-specific tissue structure, mask conserved molecular patterns, and limit the discovery of biologically relevant latent structure. Although dimension reduction is essential for representing high-dimensional molecular data in a lower-dimensional space, existing multi-slice methods typically enforce a globally shared representation that inadequately accommodates slice-level heterogeneity. To address this limitation, we propose Multi-Slice Graph Principal Component Analysis (MSGPCA), which decomposes molecular variation into shared spatial factors conserved across slices and slice-specific factors that capture local tissue microarchitecture. In downstream analyses, MSGPCA-derived representations recover spatial tissue structure, denoise molecular profiles, and reveal biologically interpretable metafeatures associated with shared and slice-specific biology. In a mass spectrometry imaging dataset comprising nonserial slices of ductal carcinoma in situ (DCIS) and invasive breast cancer (IBC), the shared factors captured broad biological differences across tissue regions, whereas the slice-specific factors revealed intratumoral spatial variation within the IBC microenvironment. In human dorsolateral prefrontal cortex ST data, MSGPCA recovered laminar cortical architecture across adjacent slices, closely aligning with expert pathologist annotations. Together, these findings demonstrate that MSGPCA resolves shared tissue architecture while preserving local microenvironmental variation in complex multi-slice spatial omics datasets.

15
Gene Program Negotiation Defines Cellular Identity in Single-Cell Transcriptomes

Sung, J.-Y.; Cheong, J.-H.

2026-07-09 bioinformatics 10.64898/2026.07.05.736629 medRxiv
Top 0.1%
26.7%
Show abstract

Single-cell transcriptomics has transformed the characterization of cellular heterogeneity by enabling systematic analysis of biological gene programs. However, existing computational approaches primarily quantify the activity of individual programs independently and therefore provide limited insight into how multiple simultaneously active programs collectively determine cellular identity. Here we present Gene Program Negotiation (GPN), a graph-based computational framework that models regulatory decision-making among concurrently active biological programs. GPN reconstructs cell-specific program interaction networks from local transcriptional neighborhoods and quantifies regulatory organization using the Gene Program Coherence Index (GPCI) together with measures of local regulatory conflict, program diversity, and dominance. These graph-derived properties enable the classification of individual cells into five regulatory decision states: Consensus, Competition, Negotiation, Dominance, and Low activity. Applying GPN to gastric cancer single-cell transcriptomes revealed that cells sharing the same dominant biological program frequently occupied distinct regulatory decision states, demonstrating that dominant program identity alone does not uniquely define cellular regulatory organization. Competition states consistently exhibited elevated local regulatory conflict and were preferentially enriched among transition-like cells, indicating that regulatory competition is closely associated with transcriptional plasticity. Independent validation using glioblastoma single-cell transcriptomes reproduced these regulatory patterns without modification of the computational framework, supporting the robustness and generalizability of the approach across biologically distinct malignancies. These findings establish regulatory negotiation as an additional layer of cellular organization beyond conventional gene-program activity analysis. By explicitly modeling interactions among simultaneously active biological programs, GPN provides a general computational framework for investigating regulatory coordination, cellular plasticity, and dynamic cell-state organization in single-cell transcriptomic data.

16
Unbalanced Perturbation Dynamics For Cell Fate Design

Peng, Q.; Wang, Y.; Li, J.; Wang, X.; Xiao, Y.; Zhou, P.

2026-07-04 bioinformatics 10.64898/2026.06.30.735555 medRxiv
Top 0.1%
26.5%
Show abstract

Large-scale single-cell perturbation sequencing provides an unprecedented opportunity to construct virtual cells for the in silico simulation of cellular responses and the inverse design of optimal interventions. However, most perturbation-response models treat cellular responses primarily as mass-preserving shifts in transcriptomic state, whereas single-cell perturbation measurements are inherently unbalanced: the recovered endpoint population is shaped by technical sampling as well as biological perturbation-induced proliferation, apoptosis and selection. Here we introduce U-Pert, an unbalanced generative framework that learns condition- and context-dependent perturbation dynamics from unpaired single-cell snapshots. U-Pert jointly models transcriptomic state transitions and cell-number dynamics, enabling scalable and robust forward prediction of unseen perturbations and contexts, as well as inverse design to screen for desired genetic or pharmacological interventions that achieve user-defined transcriptomic or population-level outcomes. Across controlled simulations, genetic perturbation benchmarks, sciPlex3 drug responses and PBMC cytokine perturbations, U-Pert predicts unseen responses, captures both molecular and abundance changes, and performs inverse design for target gene-expression programs and cell-type compositions. These results show that cell abundance is an integral component of the perturbation phenotype, providing a mass-aware framework for virtual-cell modeling and perturbation cell fate design.

17
Microenvironment-informed inference of transcriptional progression geometry

Kobara, S.; Rahman, S. A.; Ribeiro, S. P.; Coopersmith, C. M.; Kamaleswaran, R.

2026-08-24 bioinformatics 10.64898/2026.08.21.746284 medRxiv
Top 0.1%
26.3%
Show abstract

We present BIOCURRENT, a causal inference framework that reconstructs donor-specific pseudotime geometry in transcriptomic data. By modeling gene expression as a function of baseline characteristics, microenvironmental context, and latent pseudotime, BIOCURRENT enables comparison of compressed or expanded progression intervals across transcriptional state transitions. We introduce $\Delta\Delta T$, a geometry-based estimator that quantifies differences in pseudotime intervals across conditions, enabling evaluation of changes in pseudotime intervals under hypothetical modulation of microenvironmental programs. Applications to thymic T-cell developmental lineages and to COVID-19 immune dysregulation reveal condition- and donor-specific distortions of progression intervals. Counterfactual simulation links microenvironmental context to changes in specific intracellular state transition intervals. By localizing deviations in pseudotime geometry, BIOCURRENT identifies whether shifts in transcriptomic programs emerge early or later along transcriptomic coordinates and reveals upstream programs associated with these distortions. Such localization supports transcriptional stage-aware mechanistic hypotheses and suggests candidate intervention checkpoints in complex biological systems.

18
Deep Mechanistic Models reveal pathway-extrinsic drivers of mammary MAPK signalling heterogeneity

Fabrini, G.; Froehlich, F.

2026-07-31 systems biology 10.64898/2026.07.30.741759 medRxiv
Top 0.1%
25.5%
Show abstract

Cells sense and respond to their environment through signalling pathways, and the dynamics of these pathways shape cell fate even within genetically identical populations. Two largely separate computational traditions describe this behaviour: mechanistic differential-equation models and representation-learning methods. Mechanistic models encode pathway topology and kinetics but cannot easily represent variation arising outside the modelled pathway. Representation learning, instead, maps genome-wide measurements onto low-dimensional manifolds but offers no mechanistic account of how the resulting cell states execute their functions. Reconciling these views, explaining signalling heterogeneity in a manner that is at once data-driven and mechanistically interpretable, has remained difficult. Here we introduce deep mechanistic models (DMMs), which couple semi-supervised representation learning to an ordinary-differential-equation model of EGFR/MAPK signalling, trained end-to-end so that the learnt representation and mechanistic parametrisation inform each other. Applying DMMs to multiplexed signalling data from 63 breast cancer cell lines, we show that the models generalise to held-out cell lines and attribute most heterogeneity to pathway-extrinsic factors, namely baseline ERBB2 activation and a Ca{superscript 2}/p38 signalling axis, rather than to variation in core MAPK components. Where the models fail, the discrepancies pinpoint rare signalling-altering mutations and recurrent programmes, including a putative AMPK- BRAF MEK-inhibitor-resistance axis and a cytoskeletal programme. We further find that mechanistic integration of EGFR receptor levels reshapes the learnt representation, rendering a molecular and a systems-level account of the same data equivalent in predictive power. DMMs thus offer a general framework for fusing mechanism with learning, naturally extendable to further modalities such as imaging, and simultaneously turn model failure into a systematic route to discover and evaluate candidate biology.

19
Synthesizing Mechanistic Hypotheses from Single-Cell Omics via Discretized Feature Attribution and Empirical Language Model Grounding

Chen, J.; Hong, Y.; Bermudez, A.; Hu, J.; Hsieh, C.-J.; Lin, N.

2026-07-10 systems biology 10.64898/2026.07.09.737344 medRxiv
Top 0.1%
22.1%
Show abstract

Single-cell multimodal omics offer unprecedented resolution of cellular networks, yet translating continuous computational attributions into structured, testable biological mechanisms remains a persistent bottleneck. To address this limitation, we introduce an analytical pipeline employing decision trees to discretize continuous neural network attributions into explicit regulatory thresholds. These boundaries then structurally constrain large language models, enabling them to integrate established literature with empirical data to synthesize context-specific hypotheses. Applying this continuous-to-discrete framework across sparse datasets yielded novel biological mechanisms. Specifically, the framework articulated a cytoskeletal gating hierarchy governing EGF-stimulated pathways, identified transcriptomic drivers of input resistance in cortical interneurons, and delineated translational logic predicting Ki-67 abundance within spatial transcriptomics. Retrospective benchmarking validated the capacity of the framework to autonomously reconstruct published regulatory logic. Supported by a locally deployable open-weight language model and a code-free interface, this approach establishes an auditable methodology to extract robust experimental hypotheses from high-dimensional single-cell data.

20
Assay concordance sets exact ceilings on what one biological score can predict

Liu, Z.

2026-08-26 bioinformatics 10.64898/2026.08.24.746774 medRxiv
Top 0.1%
21.8%
Show abstract

Computational models of biology are ranked by averaging one prediction against many experimental realizations of a phenotype that are treated as interchangeable. We show this imposes an exact, model-free ceiling fixed by how much those realizations agree with each other, and that the ceiling depends on the evaluation metric through a single support-function identity. Measuring assay concordance across four public registries, 2,822 MaveDB score sets, 217 ProteinGym assays, two drug screens and 1,150 CRISPR cell lines, we find that two assays of one target agree at 0.56-0.68, and that 541 domains measured twice with different proteases fix assay reliability at 0.897, so 70-90% of every ceiling is irreducible biology rather than noise. Published predictors realize 63% of the achievable on the correlation benchmarks report and 18% on the top-1% selection their users perform. We provide the estimator, the ceilings, and the measurements the field has not made.